Remote Sensing in Ecology and Conservation
○ Wiley
Preprints posted in the last 90 days, ranked by how well they match Remote Sensing in Ecology and Conservation's content profile, based on 14 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Vallery, A. C.; Kabra, K.; Gibbons, R.; Arnold, H.; Minnich, N.; Barman, A.
Show abstract
Waterbirds serve as important indicators of both aquatic and terrestrial ecosystem health, making effective monitoring essential for tracking population health and identifying potential causes of decline. Drones have provided opportunities to overcome historic waterbird monitoring challenges, but the expertise and time required for manual image analysis creates a major bottleneck. Recent advances in deep learning-based object detection have enabled rapid, automatic detection of features in complex ecological imagery, though applications have largely been limited to single-species colonies, and practitioners lack quantitative comparisons of annotation time and accuracy across different levels of automation. We systematically compared four waterbird monitoring approaches using identical survey areas from Chester Island, a mixed-species colony in Matagorda Bay, Texas, in 2025: (1) traditional ground-based counts, (2) manual drone imagery-based counts, (3) computer-assisted counts using pre-annotations from an object detector with manual human verification (Human+ML), and (4) fully automated counts using object detector annotations (ML-only). We trained a YOLOv10 object detection model on manually annotated imagery of Chester Island in 2021 and applied it to the 2025 imagery. Manual drone annotation detected 6,530 birds in 40.5 hr and served as the primary reference standard. Human+ML detected 5,826 birds (89% of manual) in 7.7 hr, an 81% reduction in annotation time. ML-only detected 5,679 birds (87% of manual) in approximately 46 min, a 98% reduction. Ground counts recorded 5,868 birds (90% of manual). Detection generalized well across species while classification depended heavily on training data and morphological distinctiveness. The Human+ML workflow emerged as a practical middle ground, providing practitioners with empirical data to evaluate partial versus full automation strategies based on monitoring objectives.
Bjerge, K.; Wogram, S. F. A.; Serra-Marin, P. E.; Sakhiashvili, O.; Hoye, T. T.
Show abstract
Automated monitoring of insect pollinators in natural environments with insect camera traps and trained deep learning algorithms provides novel data for insect ecological studies. However, efficient and accurate image recognition analysis of the recorded images or videos is challenging, particularly for images containing small insects against complex backgrounds with diverse vegetation communities. Even when insects can be detected in images, identifying their taxonomy remains difficult, particularly in footage with low image resolution, light conditions, and distances from the plants, and in cases where insects appear blurry or only partially visible. In this work, we present InsectDCT, an AI-based pipeline for automated detection, hierarchical classification, and tracking of insects in footage of natural vegetation tested in different environments. The InsectDCT pipeline consists of three levels: insect Detection and localization, hierarchical taxonomic Classification, and spatio-temporal Tracking. In the first stage, insects are detected in time-lapse images or video recordings using the You Only Look Once (YOLO11) object detection architecture. Detection performance is improved using motion-enhanced images, which improve robustness in cluttered and 3 dimensional environments. The detector is trained on an extensive dataset that contains more than 60,000 images collected using camera traps deployed across a wide range of plant families and floral habitats. In the second stage, detected insects are classified using a hierarchical taxonomy-aware classification framework that covers 80 taxonomic groups. Classification is performed at multiple taxonomic levels, including order, family, and genus/species, allowing coarse and fine-grained ecological analyzes while accounting for varying levels of visual ambiguity. In the third stage, a multi-object tracking module is applied to high temporal-resolution image sequences and video data to associate detections of the same individual across time. InsectDCT code and all datasets are made publicly available. Author summaryInsects are declining worldwide, creating an urgent need for efficient methods to monitor their abundance, activity, and diversity. Traditional insect surveys often require extensive fieldwork and expert taxonomic identification, which limits the scale and frequency of monitoring. In this study, we developed InsectDCT, an artificial intelligence-based pipeline that automatically detects, classifies, and tracks insects in camera-trap recordings collected from natural and semi-natural environments. Our approach combines deep-learning methods for object detection, hierarchical taxonomic classification, and tracking of individual insect observations through time. Unlike many existing systems that are trained for a single habitat or plant species, we designed our framework using images collected across a wide range of flowering plants, camera systems, and insect groups. This makes the system more transferable to new ecological settings. The classifier can identify insects at multiple taxonomic levels and can return higher-level classifications when species-level identification is uncertain. We demonstrate that the pipeline can process large image datasets efficiently, including on low-power edge-computing devices such as Raspberry Pi systems. By providing both the software and the underlying datasets, we aim to support scalable, non-invasive insect monitoring and facilitate future ecological and conservation research.
Owens, G.; Wood, C. M.; Hunt, T. J.; Bussolini, L. T.; Kriesl, A.; Alves, F.; Stojanovic, D.
Show abstract
Efficiently finding rare species is a perennial challenge in conservation science. The orange-bellied parrot Neophema chrysogaster is a rare mobile bird that is difficult to locate using traditional field survey techniques with human observers. We harnessed recent advances in bioacoustic technology to create a survey framework that integrates passive acoustic surveys and semi-automated detection to increase monitoring capacity for the orange-bellied parrot. We developed a custom BirdNET classifier for the orange-bellied parrot and compared efficacy of acoustic and field surveys using an occupancy framework. We deployed autonomous recording units across the orange-bellied parrots breeding range in southwest Tasmania and concurrently undertook between three and six repeated point-count surveys at the same 48 sites using human observers. Our custom BirdNET classifier had high accuracy and discrimination abilities. Validation of model scores across a week (5,712 hours of audio) required 60 hours reviewing time and yielded a 95% confidence of a correct BirdNET prediction at scores over 0.998. Occupancy analysis showed that the detection probability of acoustic surveys (p = 0.80) was more than eleven times greater than field surveys by skilled ecologists familiar with the species (p = 0.07). We provide a template for how to implement monitoring of the orange-bellied parrot and recommendations for how our methods can be improved to optimise the classifier to account for other species and locations.
Pawlak, C. C.; Yost, J. M.; Ventura, J.; Guizan, G.; Arnold, S.; Okin, G. S.; Cavanuagh, K. C.; Fricker, G. A.; Ritter, M. K.; Gillespie, T.
Show abstract
Statewide tracking of urban tree canopy change is essential for evaluating progress toward policy targets, but detecting real change requires both high-resolution mapping and rigorous uncertainty estimation. We produced a four-year canopy cover time series for all California census-designated places using 60-cm NAIP aerial imagery and a U-Net deep learning model trained with semi-automated LiDAR-derived labels and manually annotated tiles. Canopy cover and change were estimated using stratified, error-adjusted area estimation, enabling comparisons across years. Statewide canopy cover showed a modest negative trend from 2016 to 2022 (Sens slope: -0.60% per year), but confidence intervals included zero across all groups and climate zones, indicating that trends were not statistically distinguishable from no change. Urban canopy cover was consistently lower than non-urban canopy by approximately six percentage points, and canopy cover was highest in the Northern California Coast and lowest in the Southwest Desert. Residential parcels accounted for 55-56% of canopy within incorporated urban areas across all years, indicating that statewide canopy increase goals will require engagement with private landowners. Error adjustment substantially altered canopy estimates relative to raw pixel-count totals, with direct implications for AB 2251 canopy tracking where baselines and targets drawn from unadjusted maps may not reflect true canopy extent. This open-source workflow is transferable to future NAIP acquisition years and other U.S. states, providing a scalable framework for long-term urban forest monitoring.
Nanduri, N.; Ogundare, J.; Anderson, G.
Show abstract
Camera trap networks such as Snapshot Safari have generated millions of labelled wildlife images across Africa, enabling the training of deep learning models for automated species classification. However, deploying models trained in one African region to another remains poorly understood. To the best of our knowledge, this study presents the first systematic evaluation of geographic domain shift within the African continent for wildlife camera trap species classification, using the Machine Learning sub-field of Artificial Intelligence. We use three model architectures, each interacting with Snapshot Serengeti in a different way: BEiTV2is fine-tuned on Serengeti images as a supervised baseline; DINOv2 with FAISS uses Serengeti images as a retrieval index without any weight updates; and BioCLIP is a true zero-shot foundation model that receives no Serengeti training data at all. All three are then evaluated on two Southern African test sets, Snapshot Kgalagadi and Snapshot Kruger, as well as on locally collected wildlife photographs from Botswana. We conduct eight experiments covering in-domain baselines, cross-dataset transfer, data scaling, MegaDetector preprocessing, grayscale vs. colour image conditions, and per-species transfer analysis. This work provides the first empirical characterisation of intra-African domain shift across both supervised and zero-shot architectures, and offers practical guidance for conservation AI practitioners who need to deploy models across the diverse ecosystems of Southern Africa without collecting new labelled data.
Gibbons, A.; Parnell, A.; Donohue, I.; Ogasawara, M.; Ross, S. R. P.-J.
Show abstract
O_LIMonitoring and limiting the spread of invasive species on islands requires efficient detection and population estimation methods. However, elusive species can be difficult to monitor using traditional methods, making autonomous approaches such as camera trapping and acoustic monitoring increasingly valuable. C_LIO_LIOn the island of Okinawa, Japan, the small Indian mongoose ( Urva auropunctata) threatens many native species since its introduction in 1910. Listed among the worlds worst invasive species, effective monitoring of U. auropunctata in Okinawa is critical. The Okinawa Environmental Observation Network (OKEON) uses camera traps to detect U. auropunctata, but success depends on precise placement. Though OKEON also includes a high-resolution acoustic monitoring programme, no audio classification model currently exists for U. auropunctata. Developing such a model could improve substantially our capacity to detect and manage the species. C_LIO_LIUsing sparse U. auropunctata vocalisations collected from camera trap videos, we built a lightweight Convolutional Neural Network distilled from a more complex model for classifying contact calls and alarm calls of U. auropunctata. Our distilled model performed similarly to the full model at detecting vocalisations from training data, but was considerably faster. C_LIO_LIWe applied the distilled classifier to [~]486 hrs of audio collected over eight years from southern Okinawa, where we successfully detected U. auropunctata a handful of times in each year of recording. In spite of strong model performance on test data, our model did not transfer well to unseen data, perhaps owing to the rarity of U. auropunctata calls and consequent small training dataset size, limiting its utility for ecological monitoring. C_LIO_LIPractical implication. The use of sparse audio data from camera trap videos to train an acoustic classifier had limited utility to detect the rarely vocalising U. auropunctata from passive acoustic monitoring data. We provide several recommendations for enhancing classifier performance to provide robust actionable insights into the distribution and spread of U. auropunctata, and aid targeted conservation efforts for Okinawas threatened biodiversity. C_LI
McMurry, S.; Alyetama, M.; Goldstein, B.; Kays, R.
Show abstract
Models for estimating animal density from camera traps require four parameters informing detection: movement speed, daily activity level, staying time (duration animals remain within the detection zone), and effective detection distance. These parameters traditionally come from labor-intensive manual measurements and auxiliary telemetry. Recent advances in computer vision can provide the positions of animals in camera trap images, which have been used for distance sampling. We extend this approach to extract all four parameters from imagery, providing the first AI-derived estimates of movement speed and staying time from automated coordinate tracking. We also introduce a new joint multi-species hierarchical distance function that estimates deployment-specific effective detection distances while borrowing strength across species through partial pooling. Our pipeline integrates MegaDetector for animal detection, the Segment Anything Model for segmentation, and Dense Prediction Transformers for monocular depth estimation. From frame-level coordinates, we reconstruct movement trajectories across burst sequences to estimate speed with size-biased distribution corrections, calculate staying time through bounding box interpolation, and estimate activity levels from detection timestamps. The joint hierarchical distance function decomposes the detection scale parameter into a shared deployment-level effect and species-specific offsets, so species effects represent deviations from the multi-species average, allowing data-rich species to inform detection conditions where rare species have few observations. AI-derived scene depth enters the model as a covariate on detection range, providing a vegetation openness metric from the same pipeline. To address position errors from depth estimation, we apply data quality filters. We processed 122,574 frames from 181 deployments across montane forests in Washington and Montana, generating parameter estimates for 12 species without manual annotation. Automated speed estimates produced day ranges 2.7 to 4.3 times GPS telemetry-derived daily distances, reflecting differences between encounter velocity within detection zones and landscape-scale displacement. Deployment-level variation in detectability exceeded species-level differences 3:1, with scene depth strongly predicting detection range; mean effective detection distances ranged from 4.1 to 7.6 m. Applied to a Random Encounter Model, these parameters yielded a white-tailed deer density estimate of 21.4 animals/km{superscript 2} and the Random Encounter Staying Time model yielded 11.6animals/km{superscript 2} in Montana. This pipeline enables scalable density estimation across large camera trap networks.
Chia, W. H.; Jahanshahi, I.; Loh, L. Y.; Zheng, A.; Verma, N.; Mussman, S.; Shi, B.; Stroud, J. T.
Show abstract
Community science platforms like iNaturalist generate unprecedented volumes of biodiversity data, but their scientific utility depends critically on accurate species identification--a persistent challenge when contributors often lack taxonomic expertise. We developed "LizardLens", a two-stage machine learning pipeline that decouples object detection from species classification to enable fine-grained identification of morphologically similar organisms in visually complex field photographs. Using 10,000 verified iNaturalist images of five Anolis lizard species in Florida, we trained specialized YOLO-based detection and Swin Transformer classification models and compared performance against state-of-the-art single-stage architectures. Our two-stage pipeline achieved 83.0% Top-1 accuracy and a macro-averaged F1-score of 89.0%, indicating strong precision-recall performance across species and outperforming single-stage YOLOv8 and YOLOv12 models across all evaluation metrics for all species, with relative improvements ranging from 10.5% to 13.2%. Gradient-weighted Class Activation Mapping (Grad-CAM) indicated that the models predictions were consistently associated with regions corresponding to diagnostic morphological (e.g., head shape, feet, and limb lengths) and pattern features (e.g., ocular rings and body patterning), providing evidence that LizardLens leverages biologically relevant visual cues consistent with those used by expert taxonomists. Error analysis identified partial occlusion and multiple proximate individuals as primary sources of missed detections, while spurious detections of lizard-like environmental features (e.g., sticks, bark) represented the dominant false positive error mode. We deployed LizardLens as an accessible web application featuring interactive bounding box correction, ranked species predictions with confidence scores, directly supporting the "Lizards on the Loose" middle school community science initiative. By combining technical advances in fine-grained visual classification with user-centered design, LizardLens demonstrates how machine learning can simultaneously enhance data quality for biodiversity monitoring and provide authentic scientific experiences for student participants. Our approach is generalizable to other small-bodied organisms in complex habitats and provides a framework for translating computer vision advances into practical tools for community science and conservation.
Remy, E.; Carlier, A.; Massol, E.; Kacimi, R.; Chaine, A. S.; Cauchoix, M.
Show abstract
Widespread arthropod declines pose risks to ecosystem functioning and agriculture. Assessing this decline or potential remediation implies the need for standardized and scalable population monitoring. Image-based methods, including camera traps and citizen science programs, are increasingly used, but the volume of data collected requires automated analysis. Robust arthropod detection is essential for individual counting or fine-grained classification, yet current datasets and algorithms do not address the vast morphological diversity across arthropod species and often overlook the variety of photographic contexts, such as differences in background, lighting, and image composition, in which arthropods are captured. To address this gap, we developed an arthropod detection dataset, covering all terrestrial families present in France with available validated images on the iNaturalist platform (749 families). To achieve this, we employed an iterative workflow in which a YOLOv11 model pre-annotated images -- using one representative species per family-- followed by manual correction and model retraining. Repeating this process progressively reduced annotation effort and improved model accuracy. The final outcome consists of a publicly available curated detection dataset and a robust arthropod detector for natural background scenes. The detector achieves an F1-score of 0.91, demonstrating strong performance despite substantial interspecific morphological variation and heterogeneity in photographic contexts. We further demonstrated the taxonomical universality of the model showing high F1-score and IoU averaged at the class (0.79, 0.85) and order level (0.82, 0.86) and also a good detection generalizability (F1-score>0.90, IoU>0.83) on species, genera and families never encountered by the model during training. Finally, we show how this model can be improved to generalize to new datasets using data augmentation, complementary training data or fine-tuning and increase detection of small objects. In particular, we report performance of the improved models on three use cases largely used in non lethal insect monitoring: (i) diurnal pollinator monitoring through citizen science or (ii) flower and nocturnal insects monitoring through smartphone time-lapse of a UV-illuminated white panel. These results mark an important step toward automated analysis of arthropod images in natural contexts, from both large-scale automated monitoring approaches or from citizen science monitoring programs.
Oliveira, M. B.; Bernardino, H. S.; Vieira, A. B.; Barroso, A. A.; Augusto, D. A.
Show abstract
The automated classification of animals from photos is important in ecology and conservation biology for organizing and understanding the immense diversity of species, as well as facilitating effective conservation and management practices. It is equally important for disease surveillance systems, allowing prompt detection of anomalies in species distributions and boosting citizen-scientist platforms by making user-reported data more accurate and convenient. Image classification uses photos and can also rely on the geographical locations of animals to improve performance. While image classification models have difficulties in classifying low-quality images, unbalanced datasets, and with a small number of images, species distribution models have difficulty in classifying species that coexist in a given region. We propose here strategies for combining image classification models based on deep neural networks with species distribution models using genetic algorithms. The proposal is applied to a real-world dataset comprising fifteen classes of animals from the Brazilian fauna obtained from Fiocruzs citizen-scientist Wildlife Health Information System (SISS-Geo). The SISS-Geo photos portray the reality of animals in their environments, with varying quality, and pose numerous difficulties for classification. Experimental results demonstrate that the proposed integration consistently outperforms standalone models. While individual SDMs achieve Top-1 accuracies of 27.79% (MaxEnt) and 31.76% (Bioclim), and CNN-based classifiers reach 58.17% with ResNet50 and 64.13% with ResNet-152, the hybrid strategies yield substantial improvements. The genetic algorithm-based integration with a single global weight achieves up to 67.96% Top-1 accuracy, whereas the class-specific integration using fifteen parameters attains the best overall performance, reaching 69.03%.
Lesmeister, D. B.; Jenkins, J.; Ruff, Z.; Rugg, N.; Abidari, S.; Beavers, S.; Christophersen, R.; Davis, R. J.; Terwilliger, M.; Henderson, A.; Henson, B.; Kasper, J.; McCafferty, C.; Press, D.; Reffler, S.; Swingle, J.; Thomas, A.; Wert, K.
Show abstract
The Northwest Forest Plan (NWFP) Passive Acoustic Monitoring (PAM) program is a large-scale interagency biodiversity monitoring framework designed to assess the status and trends of northern spotted owls (Strix occidentalis caurina), barred owls (Strix varia), marbled murrelets (Brachyramphus marmoratus), and broader forest biodiversity across federally administered lands in the Pacific Northwest. In 2025, we deployed autonomous recording units at 2,095 sampling stations within 532 5-km2 hexagons randomly selected across approximately 24 million acres of federal forest lands in Washington, Oregon, and California. These deployments generated 1.37 million hours of acoustic recordings that we processed using the convolutional neural network PNW-Cnet v5 for automated species identification and subsequent human validation of focal species detections. Northern spotted owls were detected in 28% of sampled hexagons range-wide, including 16% in Washington, 25% in Oregon, and 60% in California. Occupancy patterns remained consistent compared to previous years, with higher occupancy concentrated in southern portions of the geographic range and continued low occupancy across much of the Washington Cascades and Oregon Coast Range. Barred owls were detected in 86% of sampled hexagons and remained broadly distributed throughout most of the NWFP area. Marbled murrelets were detected in 50% of reviewed hexagons within NWFP marbled murrelet management zones, with highest occupancy occurring in coastal forests of Oregon and Washington. The 2025 field season occurred under substantial operational constraints that reduced sampling effort by approximately half relative to 2023 and 2024 because of staffing limitations affecting participating federal agencies. Despite these reductions, the NWFP-PAM framework continued to provide broad-scale, spatially representative ecological information across the NWFP area. Results highlight the growing importance of passive acoustic monitoring and machine learning approaches for long-term biodiversity monitoring under changing environmental and operational conditions.
Eddington, V. M.; Fradet, D. T.; Craig, E. C.; Cimino, M. A.; White, E. R.; Kloepper, L. N.
Show abstract
Migratory seabirds are valuable indicators of marine ecosystem change but can be difficult to monitor during the breeding season due to dense colonies, remote breeding sites, and sensitivity to investigator disturbance. Passive acoustic monitoring offers a minimally invasive alternative to traditional surveys; however, high call overlap in large colonies complicates approaches that rely on identifying individual vocalizations. In this study, we evaluate acoustic energy as a simple soundscape metric for monitoring breeding phenology in colonial seabirds. Using a comparative approach, we deployed autonomous recorders at breeding colonies of Adelie penguins (Pygoscelis adeliae) in the Western Antarctic Peninsula and common terns (Sterna hirundo) in the Gulf of Maine. We examined seasonal patterns in acoustic energy and compared these trends with known breeding stages and colony observations. Across both species, acoustic energy exhibited distinct seasonal patterns that correspond to key phenological stages, including courtship, incubation, chick rearing, and fledging. These stages are associated with distinctive colony-wide behavioral shifts in colony attendance, territorial interactions, and parent-offspring communication that structure the breeding-season soundscape. Our results demonstrate that colony-wide acoustic energy can capture key phenological transitions in seabird colonies and provide a scalable, minimally invasive approach for monitoring breeding dynamics in remote or rapidly changing environments. HighlightsO_LIPassive acoustic monitoring can track bioindicator phenology under climate change C_LIO_LIAmplitude captures colony-level activity in dense seabird colonies C_LIO_LISoundscape patterns correspond to key breeding stages C_LIO_LIEffective in both temperate and polar seabird systems C_LIO_LIEnables scalable, low-disturbance monitoring in remote systems C_LI
Sheldon, D.; Winner, K.; Deznabi, I.; Bernstein, G.; Bhambhani, P.; Lin, T.-Y.; Desmet, P.; Dokter, A. M.; Horton, K. G.; Nilsson, C.; Van Doren, B. M.; Farnsworth, A.; La Sorte, F. A.; Maji, S.
Show abstract
The US NEXRAD radar network has monitored the aerosphere over the US and its territories continuously since the 1990s and archived nearly 300 million radar volume scans. These data contain a wealth of information about the movements of birds, bats, and insects. Historically, this biological information was difficult to access due to the amount of data and challenges in analyzing it. In the last 15 years, fueled by computational and methodological advances, large-scale aeroecology research has blossomed. However, comprehensive analyses of the NEXRAD archive remain very costly. We collected measurements of biological activity from every volume scan in the NEXRAD archive--nearly 300 million data files total--to assemble a dataset of aerial biomass over the US from 1995 to 2025. The core data are vertical profiles, which summarize biological activity at different heights above the radar station for each volume scan. We also provide time series data products that aggregate vertical profiles to point measurements at radar stations across time. These data products can support a range of aeroecology analyses at significantly reduced effort.
Perez-Granados, C.; Morant, J.; Funosas, D.; Sebastian-Gonzales, E.
Show abstract
Recent advances in automated technologies, such as passive acoustic monitoring, provide a powerful framework for surveying bird communities at broad spatial scales. Among the most widely used artificial intelligence tools for automated bird sound recognition is BirdNET, which can identify over 6,000 species worldwide. However, the effects of key user-defined settings, such as species filtering, remain poorly evaluated. Here, we assess how alternative species-filtering strategies influence BirdNET performance in describing bird communities worldwide. We analysed 5,047 minutes of sound recordings from 72 locations worldwide, comprising 1,192 bird species identified by expert ornithologists. We compared three common species-filtering approaches applied in BirdNET workflows to post-process its output: no filtering, spatial filtering (species present all-year at a given location), and spatio-temporal filtering (species present at a given location and week). The unfiltered approach maximised BirdNET species detection (recall) but suffered very low precision (had many misidentifications) and poor overall performance. In contrast, the other two filtering strategies greatly improved precision and overall performance, despite moderate reductions in recall. Among them, spatio-temporal filtering consistently achieved the best performance across most datasets and regions globally. Within this optimal filtering approach, we also evaluated the role of another parameter: occurrence probability thresholds. Intermediate values of this threshold (around 0.05) maximized BirdNET performance in community-level analyses. Our results demonstrate that species filtering is a key but often underappreciated component of BirdNET workflows. We hope our findings may guide future studies in selecting optimal species filters, while emphasising that filtering selection should be guided by study objectives and data context.
Ardila-Villamizar, M.; De Clippele, L. H.; Dominoni, D. M.
Show abstract
Convolutional Neural Networks (CNNs) have become increasingly prominent in biodiversity monitoring due to their strong performance in accurately detecting species from sound recordings, overcoming some limitations of traditional methods such as point-counts. Yet, their use in urban ecosystems remains limited, highlighting the need for frameworks that identify modelling strategies to optimize their performance in these complex soundscapes. Here, we evaluated how preprocessing and labelling strategies, detection thresholds, sample size, and architecture affect the performance of CNNs for bird identification in urban tropical ecosystems. We also assessed its potential by comparing CNN-derived biodiversity estimates with those from point-counts and acoustic indices. For this, we used one week of recordings collected along urbanization gradients in five Colombian Andes cities to developed 11 multiclass CNN models varying in spectral representation, labelling strategies, training data source and backbone architecture. The best-performing model, evaluated with F1-scores, combined Log-Mel spectrograms, multispecies labels, ecosystem-specific recordings, a probability threshold of 0.3 and a ConvNeXt backbone with its performance generally improving with sample size. Although CNNs and point counts detected partially distinct assemblages, CNN-derived species richness was comparable to that estimated from point-counts. In addition, the Normalized Difference Soundscape Index (NDSI) was positively associated with richness, suggesting its potential as a biodiversity proxy in tropical urban soundscapes. Overall, by identifying effective modelling designs and monitoring strategies, our study advances the development of robust biodiversity assessment frameworks in urbanized ecosystems in the Neotropics whilst also providing methodological guidance for future research and practical insights for wildlife monitoring and conservation.
Hayati, A.; Gong, J.; Nagesh, V.; Avci, P.; Ong, A. Y.; Masalkhi, M.; Engelmann, J.; Karouia, F.; Scott, R. T.; Keane, P. A.; Costes, S. V.; Sanders, L. M.
Show abstract
Space-biology imaging studies are often constrained by severe data scarcity, limiting the development of robust machine-learning biomarkers. Rodent spaceflight and space-analog datasets provide an important preclinical setting for testing transfer-learning strategies, but the extent to which human retinal foundation models can generalize to rodent optical coherence tomography (OCT) remains unclear. Here, we benchmark cross-species adaptation of RETFound, a human retinal Vision Transformer pretrained on 1.6 million retinal images, for chronological age prediction from Brown Norway rat OCT B-scans in the NASA Open Science Data Repository dataset OSD-679. We adapted RETFound using Low-Rank Adaptation (LoRA) and evaluated performance on control animals under matched 3-fold rat-level cross-validation. We compared RETFound+LoRA with a strong ImageNet-pretrained Xception baseline under matched protocols and included a scratch/random ViT as a supplementary negative-control architecture check. Metrics included mean absolute error (MAE), R2, and inter-eye mean absolute difference (MAD). RETFound+LoRA achieved MAE = 26.20 {+/-} 5.03 days with R2 = 0.744 {+/-} 0.049. However, Xception performed better in the primary benchmark (MAE = 19.01 {+/-} 7.67 days, R2 = 0.853 {+/-} 0.082), and the matched-fold comparison favored Xception, although this result should be interpreted cautiously given the small number of folds. Inter-eye consistency was maintained across the matched control evaluation, and saliency maps localized model attention to anatomically plausible inner retinal regions. Together, these results show that human retinal foundation models can transfer to rodent OCT in a scientifically useful way, but also that strong CNN baselines may outperform transformer-based models in small-sample cross-species settings. This preprint provides a reproducible benchmark and baseline framework for future retinal biomarker development in space biology. Significance StatementSpace-biology imaging studies are intrinsically data-limited. This preprint provides a reproducible cross-species benchmark for adapting Earth-trained retinal models to rodent OCT under small-sample conditions, highlights the value of strong CNN baselines, and offers a reusable starting point for future retinal biomarker development in space-relevant datasets.
McInnes, J. C.; Burgess, T.; Mergard, G.; Wells, M. R.; McMahon, C. R.; Neave, M. J.; Polanowski, A.; Terauds, A.; Tornos, J.; Lejeune, M.; Briand, F.-X.; Baele, G.; Boulinier, T.; Achurch, H.; Alderman, R.; Lashko, A.; Wienecke, B.; Wynen, L. P.; Viola, B.; Virtue, P.; Hodgson, J. C.
Show abstract
High pathogenicity avian influenza (HPAI) has spread across the sub-Antarctic, causing significant wildlife impacts. We report its first detection in an Australian external territory, Heard Island and McDonald Islands, which supports over one million breeding seabirds and seals. Drone and ground surveys (October 2025, January 2026), combined with viral genome analysis, confirmed infection with Influenza A H5N1 clade 2.3.4.4b at Heard Island. Drone surveys revealed mass mortality in southern elephant seals, with 8,573 pups (62%) recorded dead across Heard Island by the final surveys. Mortality increased at an average rate of 5.6% per day in a subset of harems, and the highest observed mortality in a harem was 97%. Based on the average (76%) mortality in the final surveys, total estimated pup mortality at Heard Island was 13,359 (from a total population of 17,364 pups), though this may be an underestimate as mortality was ongoing at this time. HPAI was detected in six of nine species tested and, we suspect, led to elevated mortality in king and gentoo penguins. Phylogenetic analysis indicates the virus was introduced from Crozet Islands, with an estimated arrival around August 2025. These data show the continued easterly spread of HPAI around the sub-Antarctic, with severe but heterogeneous impacts across taxa. Our results demonstrate the value of drones for large-scale monitoring, underscoring the need for continued and enhanced HPAI surveillance across the Southern Ocean.
Hendrikx, H.; Belaud, E.; Postic, F.; Scalabrino, M.; Lebeau, M.; Le Maire, G.; Jourdan, C.; Gallet, P.; Hedde, M.
Show abstract
1 - Automated in situ sensors - e.g., buried scanners - are transforming biodiversity monitoring by generating data at spatio-temporal resolutions unattainable through traditional sampling, including in cryptic environments such as soil that have remained largely inaccessible to existing methods. However, extracting ecologically meaningful information from these data streams requires substantial image processing effort that currently constitutes a critical bottleneck, particularly when the signal-to-noise ratio is low and annotated training data are scarce. 2 - Standard end-to-end deep learning detection pipelines offer unsatisfactory results due to the lack of training data and heterogeneity of the taxa of interest. We explore the potential of combining traditional computer vision algorithms with state-of-the-art deep learning models to build an efficient raw data processing pipelines from limited annotation effort. Specifically, based on the observation that the background barely changes, we focus on the differences between two consecutive images to turn the initial detection problem (with very low signal) into a simpler classification problem, which we solve by fine-tuning foundation models on limited annotated data. 3 - Our approach significantly reduces the annotation effort, allowing us to release an open dataset with about 600 soil scans and more than 8 000 labeled invertebrate occurrences across nine taxa. Using this dataset to train our models, we obtained population count estimates with relative errors ranging from 10% to 61% across taxa over a three-month period. Ecological validation through a land-use stability analysis showed full directional congruence between automated and expert-annotated classifications across all nine taxa examined, with effect-size discrepancies proportional to per-taxon classification accuracy. 4 - These results demonstrate that combining domain-specific heuristics with fine-tuned foundation models provides an effective and data-efficient strategy for automating ecological image processing workflows in low-signal, data-scarce contexts. The validated pipeline removes the manual annotation bottleneck that has historically limited scanner-based soil monitoring to short observational windows and restricted taxonomic scope, opening the way for continuous, large-scale tracking of soil invertebrate community dynamics at resolutions previously unachievable.
Potter, S.; Jansen, J.; Hill, N.; Lucieer, V.
Show abstract
Antarctic benthic organisms are highly diverse and play a critical role in the Southern Ocean ecosystem. Despite decades of sampling, vast areas of the Antarctic continental shelf remain biologically unsurveyed due to logistical and financial constraints, limiting baseline knowledge essential for effective conservation planning. Species distribution models (SDMs) allow biodiversity to be inferred in the absence of biological data by linking benthic community patterns to environmental predictors. However, the resolution of the environmental predictors, particularly bathymetry, varies significantly between regions, casting doubt about how reliably SDMs can be used to predict into regions where only coarse-resolution data are available. Here, we show that SDMs trained on high-resolution data underestimate Antarctic benthic morphospecies richness by up to 18% when applied to aggregated coarse-resolution environmental data (and up to 50% when using satellite-derived ETOPO bathymetry). Using six systematically degraded versions of high-resolution multibeam bathymetry and annotated seafloor imagery across three Antarctic regions, we evaluate SDM performance both with and without additional environmental variables. High-resolution bathymetry captures terrain complexity most effectively, but we find that the spatial distribution of richness hotspots and the median richness per cell remain consistent, provided models are applied at the same resolution at which they were trained. Our results suggest that while high-resolution bathymetry may enhance local predictions, coarse-resolution data may be more robust for regional-scale predictions, such as those used for Antarctic shelf-wide spatial planning.
Gallego, J.; Martinez-Vargas, J. D.; Lopez, J. D.
Show abstract
O_LIIdentifying individual animals from vocalizations is an emerging research area in computational bioacoustics. This non-invasive approach reduces reliance on physical capture and tagging for wildlife monitoring. Recent advances in this area leverage deep learning and bioacoustic foundation models adapted from species-level classifiers. However, these models typically rely on fixed-window inputs and may not fully capture temporal structure across extended or complex songs, which can contain information relevant to individual discrimination. C_LIO_LIHere, we evaluate whether modeling sequences of pretrained bioacoustic embeddings improves acoustic individual identification. We developed a framework that integrates transfer-learned spectrotemporal representations from BirdNET with a lightweight long short-term memory network. Unlike static baselines that use either the first embedding or an average over all embeddings, our approach processes time-ordered embedding sequences, allowing the classifier to use information distributed across multiple windows. C_LIO_LIWe evaluated the system using 87,865 vocalizations from 352 individuals across seven species. We used four publicly available vocalization datasets with individual-level labels, covering durations from 0.76 s (short calls) to 27.6 s (prolonged songs). Across five random seeds with stratified partitions, the framework achieved mean test accuracies between 93.9% and 98.3% and macro-F1 scores ranging from 93.2% to 98.3%, without data augmentation. The clearest gains over the strongest static baseline were observed for the great tit and the great spotted kiwi, reaching +1.3 and +2.8 percentage points, respectively. C_LIO_LIOur results indicate that the contribution of temporal sequence modeling depends on vocalization structure, rather than providing uniform evidence that chronological order drives performance. The benefit of recurrent aggregation was limited or absent for short calls but evident for long vocalizations with multiple informative windows. By combining bioacoustic foundation models with lightweight recurrent modeling, this approach provides a scalable, CPU-efficient tool for autonomous wildlife monitoring, particularly for species with extended or structurally complex vocalizations. C_LI